Papers with CIDEroptimized model
Fine-grained Image Captioning with CLIP Reward (2022.findings-naacl)
Copied to clipboard
| Challenge: | Modern image captioning models are usually trained with text similarity objectives . reference captions often describe only the most salient objects in images . |
| Approach: | They propose to use CLIP to calculate multi-modal similarity and use it as a reward function . they propose a simple finetuning strategy to improve grammar that does not require extra text annotation. |
| Outcome: | The proposed model generates more distinctive captions than the CIDEroptimized model on text-to-image retrieval and fineCapEval. |